01 The Big Picture
In doc 08 the enemy was bandwidth at 3,000 GB/s of HBM. On your phone the same wall exists — at ~100 GB/s of LPDDR. The physics is identical; the budget is not.
A datacenter H100 reads model weights at ~3,350 GB/s and pays for power after the fact, from a wall socket. A phone reads them at roughly 50–100 GB/s from LPDDR5, and pays during the fact, from a battery, in joules the user can feel as heat. That is the entire design space in one sentence: everything on-device is the datacenter problem with 30× less bandwidth and a power envelope of watts, not kilowatts.
This is why the on-device stack is not "a smaller cloud stack." It is an inverted set of priorities: the cloud optimizes throughput per GPU-dollar; the device optimizes joules per token and latency per tap. Models that look trivially small by datacenter standards (7B parameters) are, on a phone, a 3.5 GB weight file that must stream through a narrow memory pipe for every single token.
02 What It Is — The Edge Inference Envelope
On-device inference is running the full forward pass — weights, KV cache, sampling — on the user's own silicon: phone NPU, laptop CPU/GPU, or a browser. The defining discipline is the envelope: the fixed set of resources the model must live inside.
Memory envelope
iOS gives an app a few GB; Android often less. Weights + KV cache + activations + the rest of your app must all fit. 4-bit 7B ≈ 3.5 GB of weights alone — that is the model budget for most phones.
Bandwidth envelope
LPDDR ~50–100 GB/s vs HBM ~3,000 GB/s. Since decode speed is bytes-per-second over bytes-per-token, this single number sets your tokens/s ceiling before any software runs.
Power envelope
Sustained NPU budgets are ~1–5 W. Sustained load throttles; the phone gets warm and the OS will slow you down. You design for joules per inference, not tokens per second per dollar.
The crucial reframing for engineers: the envelope is bytes, not FLOPs. Every architectural choice on-device — quantization, GQA, small context windows, weight caching in SRAM — is a strategy for reading fewer bytes per token. Compute is nearly free at these scales; memory traffic is the whole bill.
03 Why It Exists
Given how good cloud APIs are, why compress models into a phone at all? Four pressures, none of which the cloud can ever fully answer:
04 How It Works — One Token, From Battery to Byte
Follow a single decode step through the device's memory hierarchy — then watch what "send to cloud instead" would have cost.
Read the loop again as a system: every token forces a full sweep of the weights — LPDDR → SRAM → NPU → logits — then the whole cycle repeats. The NPU is not the bottleneck; the pipe is. Which is why the next two sections are almost entirely about shrinking the pipe: fewer bytes per parameter (quantization), fewer parameters per token (distillation, sparsity), fewer bytes per attention step (GQA, small context — doc 22).
05 The Memory Math
Three formulas decide the entire on-device design space. All three are byte-counting; none involve FLOPs.
1 · Weight footprint
This is why quantization is not an optimization on-device — it is admission criteria. At 16-bit the same model is 14 GB and simply cannot ship.
2 · KV cache footprint
That "2" is K and V; n_kv = 4 instead of 32 is Grouped-Query Attention — 8× less cache than classic MHA. Without GQA the same context costs 2 GB: a second model's worth of RAM. Phones need GQA + short contexts for the same reason datacenters do (docs 08, 22) — the KV cache is read on every decode step too.
3 · Decode speed
Instant corollary: dropping weights from 8-bit to 4-bit halves bytes-per-token and doubles decode speed — quantization is a latency feature, not just a memory one, because decode is memory-bound (doc 08's law, restated in doc 10).
| Quantity | Formula | 7B @ 4-bit on phone | Same model, datacenter |
|---|---|---|---|
| Weights | N_bits·P/8 | 3.5 GB | 3.5 GB (of 80 GB HBM) |
| KV cache @ 4k ctx | 2·s·L·n_kv·d·2 | 0.25 GB (GQA) | 0.25 GB (2 GB w/o GQA) |
| Memory BW | — | ~100 GB/s | ~3,350 GB/s |
| Decode ceiling | BW / bytes-per-token | ~28 tok/s | ~957 tok/s |
| Power source | bytes × pJ/byte | battery, ~1–3 W sustained | wall socket, ~700 W |
06 Quantization & Distillation — The Tooling That Shrinks P
Sections 05's formulas have two levers: N_bits (quantization) and P (distillation/pruning). These are the toolchains that make a phone-sized envelope viable at all.
Per-channel quantization
Storing one scale per output channel works for activations and matrices, but LLM weight distributions have outlier channels that poison whole rows. Group-wise scales (every 32 or 64 weights) localize the damage. NF4 (NormalFloat4) goes further: instead of uniformly spaced 4-bit levels, it uses quantile-spaced levels of a normal distribution — the values real weights actually take — so each code is equally likely and information-per-bit is maximal.
GPTQ / AWQ — error-aware rounding
Naive rounding quantizes each weight independently. GPTQ and AWQ instead ask: what matters is the layer's output. The objective is to minimize the output error, weighted by the activation statistics of real data:
GPTQ compensates rounding errors of one weight by adjusting its not-yet-quantized neighbors, column by column. AWQ observes that a few salient channels dominate the loss, protects their precision, and pushes more aggressive bits onto the rest. Result: 4-bit models within a point or two of fp16 on benchmarks — the difference between "the math says 3.5 GB" and "the model still works at 3.5 GB."
Structured sparsity — 2:4 on tensor cores
Sparsity is quantization's sibling: another way to trade a little quality for fewer bytes and fewer operations — and unlike unstructured pruning, the "2 of 4" pattern is regular enough for hardware to exploit without gather/scatter logic.
Distillation — shrinking P itself
The student learns from the teacher's softened distribution. With temperature T > 1, probabilities sharpen less: instead of "next token: 99.9% 'the'," the teacher reveals "…and 0.02% 'a', 0.01% 'my'" — the dark knowledge of how tokens relate. That inter-class structure is worth thousands of hard-label examples, and it's why a 1–3B student distilled from a 70B teacher punches far above its parameter count. Distillation and quantization compose: distill to 3B, quantize to 4-bit → 0.75 GB, comfortably on-phone.
Quantize with calibration data from your domain; keep embeddings/lm_head at higher precision; use group size 32–64 for 4-bit; distill before quantizing; measure quality on your tasks, not MMLU alone.
Go below 3–4 bits on small models without a proven recipe (quality collapses non-linearly); quantize KV cache below 8-bit blindly (doc 22); assume fp16 benchmarks predict INT4 behavior; prune unstructured and hope hardware notices.
07 Hybrid Architectures — Edge Draft, Cloud Verify
Pure-local and pure-cloud are the endpoints of a spectrum; shipping products live in the middle. Two patterns dominate.
Speculative edge→cloud
The on-device small model drafts the next several tokens instantly (free, private, offline); the cloud model verifies them in one parallel prefill pass — exactly the draft/verify split of doc 24's speculative decoding, stretched across the network boundary. Acceptance costs the cloud ~1 forward pass for k drafted tokens; rejection falls back to cloud-native decoding. The device pays zero cloud latency for accepted spans, and drafts never containing sensitive text needn't be sent at all.
Router: who answers this prompt?
A small π_router — often the on-device model itself, or a classifier — grades each request by expected difficulty and routes:
Simple queries ("is this email rude?") stay local: q_local ≈ q_api, c_local ≈ 0. Frontier reasoning goes to the API. The router is a mixture-of-tokens at the product scale — the same "spend cheap tokens on easy work" logic as MoE's expert routing, applied to the edge–cloud boundary. (MoE itself is the next doc's subject.)
08 Engineering Takeaways — Joules and Thermals
On the device, the unit of cost is the joule. Power per inference decomposes the same way speed did — into memory traffic:
That 100× gap is why NPU architects obsess over dataflow: keep the current layer's weights and KV window resident in SRAM and reuse them across timesteps; touch LPDDR as few times as possible. A model that fits entirely in SRAM runs orders of magnitude cheaper per token than one streaming from DRAM — but SRAM is measured in MB, not GB, so most LLMs live in the DRAM-streaming regime and must "get few external reads" via aggressive batching of reuse. Apple and other NPUs pair ~60 TOPS of compute with exactly this SRAM-first design — yet the LPDDR ceiling of ~100 GB/s still caps end-to-end decode, per section 05.
09 Mental Models
Local inference is SQLite: small, private, instant, always there, bounded by your hardware. The API is a managed warehouse: vast, elastic, metered per query, and reachable only over the network. Products, like apps, usually need both. Lets you reason about: why "local vs cloud" is a storage-engine choice, not a religion.
HBM is a firehose feeding the datacenter GPU; LPDDR is a straw feeding the NPU. The model is the same liquid — the only question is how fast you can drink. Quantization is thinning the liquid (fewer bits per drop); GQA is a smaller mouthful per sip. Lets you reason about: why every on-device technique is "drink less," never "sip harder."
Anything on your desk (SRAM) is instant to consult but the desk is tiny; the library (LPDDR) is huge but every trip costs a walk. Good NPU code is a good researcher: bring a chapter to the desk, work through it thoroughly, minimize library trips. Lets you reason about: the 100× energy gap and why dataflow matters more than TOPS.
10 Common Misconceptions
"More TOPS = faster LLM." Decode is memory-bound; a 60-TOPS NPU streaming 3.5 GB through a 100 GB/s straw still lands at ~28 tok/s. TOPS matter for prefill and CNNs, not token generation.
"Quantization just loses a little accuracy." Done right (GPTQ/AWQ, group-wise, calibrated), 4-bit costs a point or two. Done naively (per-tensor absmax, no calibration), it can destroy a model. The variance is in the method, not the bit count.
"On-device means the model is 'as smart as' a small cloud model." Parameter count transfers only with equal context, sampling, and tooling. On-device models run shorter contexts (KV math), smaller output budgets, and no retrieval — judge them on the envelope they actually run in.
"Everything should move on-device for privacy." Routing is per-task. A request with zero sensitive content gains nothing from local execution but pays in quality. The right primitive is the router with a privacy term in its utility function, not blanket local-first.
"Browsers can't run real models." WebGPU gives the tab near-native access to the GPU; WASM+SIMD covers CPU fallback. Multi-megabyte GGUF-style weight formats stream and run in-tab — the browser is now a legitimate inference target, with zero install and the sandbox as the privacy boundary.